Papers by Jane Arleth Dela Cruz
Evaluating Large Language Models for Confidence-based Check Set Selection (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models have shown promise in automating high-labor data tasks, but their tendency to answer despite uncertainty and their difficulty handling long input contexts robustly are key challenges for adoption. |
| Approach: | They propose to use LLMs to prioritize information needing human judgment to identify low-confidence outputs for human review through "check set selection" using social media monitoring, they define the "check sets" as a list of tweets escalated to the disaster manager when the LLM has the least confidence. |
| Outcome: | The proposed approach outperforms random-sample check set selection in disaster tweet classification. |